Papers with natural language processing methods

10 papers
T-Know: a Knowledge Graph-based Question Answering and Infor-mation Retrieval System for Traditional Chinese Medicine (C18-2)

Copied to clipboard

Challenge: Traditional Chinese Medicine (TCM) is one of precious intangible cultural heritages of the Chinese nation.
Approach: They propose to use authorized and anonymized clinical records, medicine clinical guidelines, teaching materials, classic medical books, academic publications, etc. as data resources to build a TCM knowledge graph.
Outcome: The proposed system extracts triples from free texts to build a TCM knowledge graph.
EfficientOCR: An Extensible, Open-Source Package for Efficiently Digitizing World Knowledge (2023.emnlp-demo)

Copied to clipboard

Challenge: Existing OCR engines fail to provide accurate, cost-effective and sample-efficient character recognition for public domain documents.
Approach: EffOCR is an open-source optical character recognition package that is accurate, cheap to deploy and sample efficient to customize to novel collections, languages, and character sets.
Outcome: EffOCR model trains character retrieval problem and scales to novel collections, languages, and character sets.
Casting Light on Invisible Cities: Computationally Engaging with Literary Criticism (N19-1)

Copied to clipboard

Challenge: Literary critics often attempt to uncover meaning in a single work of literature through careful reading and analysis.
Approach: They propose to use a literary theory to analyze Italo Calvino's novel Invisible Cities to leverage contextualized representations to embed each city's description and use unsupervised methods to cluster embeddings.
Outcome: The proposed method can be applied to Italo Calvino’s novel Invisible Cities . authors compare results to similarity judgments generated by human readers .
Extracting Chemical-Protein Interactions via Calibrated Deep Neural Network and Self-training (2020.findings-emnlp)

Copied to clipboard

Challenge: Several natural language processing methods have been used to extract interactions between chemicals and proteins from biomedical text data.
Approach: They propose a method to extract chemical–protein interactions from biomedical text data . they use a pre-trained language-understanding model and calibration techniques to estimate uncertainty .
Outcome: The proposed approach achieves state-of-the-art performance on the Biocreative VI ChemProt task while preserving higher calibration abilities.
Diversity, Density, and Homogeneity: Quantitative Characteristic Metrics for Text Collections (2020.lrec-1)

Copied to clipboard

Challenge: Existing descriptive statistics are inadequate to summarize text collections by quantitative measures.
Approach: They propose a set of characteristic metrics that quantitatively measure the dispersion, sparsity, and uniformity of a text collection.
Outcome: The proposed metrics are highly correlated with text classification performance of a renowned model, which could inspire future applications.
Summarizing Patients’ Problems from Hospital Progress Notes Using Pre-trained Sequence-to-Sequence Models (2022.coling-1)

Copied to clipboard

Challenge: Problem list summarization requires a model to understand, abstract, and generate clinical documentation.
Approach: They propose a task that summarises patients' main problems from daily progress notes using input from the provider's progress notes during hospitalization.
Outcome: The proposed model outperforms two state-of-the-art seq2seq transformer architectures in summarizing patients' main problems from daily progress notes in the medical information mart for Intensive Care (MIMIC)-III.
Towards Debiasing Sentence Representations (2020.acl-main)

Copied to clipboard

Challenge: Recent work has shown word-level embeddings reflect and propagate social biases present in training corpora.
Approach: They propose a method to debias word embeddings to reduce biases at sentence level . they hope their work will inspire future research on characterizing and removing biase .
Outcome: The proposed method reduces biases and preserves performance on downstream tasks such as sentiment analysis and natural language understanding.
Constructing Word-Context-Coupled Space Aligned with Associative Knowledge Relations for Interpretable Language Modeling (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to train language models have limitations in interpretability . a Word-Context-Coupled Space (W2CSpace) is proposed to improve the performance of pre-trained models .
Approach: They propose a Word-Context-Coupled Space to replace pre-trained models with interpretable statistical logic.
Outcome: The proposed language model can achieve better performance and highly credible interpretability compared to state-of-the-art methods.
Three Real-World Datasets and Neural Computational Models for Classification Tasks in Patent Landscaping (2022.emnlp-main)

Copied to clipboard

Challenge: Patent Landscaping is one of the central tasks of intellectual property management and involves selecting and grouping patents according to user-defined technical or application-oriented criteria.
Approach: They propose to use a novel model that takes into account textual information from the patents’ full texts as well as embeddings created based on the patent’s CPC labels.
Outcome: The proposed model takes into account textual information from the patents’ full texts as well as embeddings created based on the patent’s CPC labels.
A new European Portuguese corpus for the study of Psychosis through speech analysis (2022.lrec-1)

Copied to clipboard

Challenge: Psychosis is a clinical syndrome characterized by symptoms such as hallucinations, delusions, thought disorders and disorganized speech.
Approach: They describe the creation of the first European Portuguese corpus for the identification of the presence of speech characteristics of psychosis.
Outcome: The results show that spontaneous speech presents more identifiable characteristics than read speech to differentiate healthy and patients diagnosed with psychosis.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations